Papers with hallucination management
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks. |
| Approach: | They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement. |
| Outcome: | The findings highlight the future directions in medical reasoning, physical system integration, and training simulations. |
Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing (2025.emnlp-main)
Copied to clipboard
| Challenge: | a recent study shows that benchmarking fails to explain why models outperform others . open-weight large language models have transformed the AI landscape . |
| Approach: | They use model diffing to analyze capability differences between Gemma-2-9b-it and SimPO-enhanced variants. |
| Outcome: | The proposed model diffing approach can provide fine-grained insights beyond leaderboard metrics . it can also help to identify model performance gaps, the authors say . |